Papers with labeling section titles

    1 papers
    STAPI: An Automatic Scraper for Extracting Iterative Title-Text Structure from Web Documents (2022.lrec-1)

    Copied to clipboard

    Challenge: Formal documents are organized into sections of text, each with a title . but there is no corpus of web documents annotated with titles and prose texts . cnn.com's john mccarthy and daniel mclears are working on a new title-text dataset .
    Approach: They propose a first title-text dataset on web documents that incorporates a wide variety of domains to facilitate downstream training.
    Outcome: The proposed system outperforms baseline models in terms of title-text identification.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations